Papers with valuable tool

27 papers
LingConv: An Interactive Toolkit for Controlled Paraphrase Generation with Linguistic Attribute Control (2025.emnlp-demos)

Copied to clipboard

Challenge: LINGCONV is an interactive toolkit for controllable text generation . it allows fine-grained control over 40 specific linguistic attributes spanning lexical, syntactic, and discourse dimensions.
Approach: They propose a toolkit for paraphrase generation that allows finegrained control over 40 specific linguistic attributes.
Outcome: The toolkit is available at https://mohdelgaar-lingconv.hf.space, with a demo video at https:youtu.be/wRBJEJ6EALQ.
FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation (2025.acl-demo)

Copied to clipboard

Challenge: FlagEvalMM is an evaluation framework designed to assess multimodal models . it is designed to be used for vision-language understanding and generation tasks .
Approach: They propose an evaluation framework that decouples model inference from evaluation through an independent evaluation service.
Outcome: The evaluation framework offers accurate and efficient insights into model strengths and limitations.
Modeling Aspect Sentiment Coherency via Local Sentiment Aggregation (2024.findings-eacl)

Copied to clipboard

Challenge: Existing studies have not explored aspect sentiment coherency, including its implications in adversarial defense.
Approach: They propose a local sentiment aggregation paradigm that models aspect sentiment coherency . they demonstrate the capability of LSA in adversarial defense .
Outcome: The proposed model outperforms existing models and achieves state-of-the-art sentiment classification performance.
Concept Over Time Analysis: Unveiling Temporal Patterns for Qualitative Data Analysis (2024.naacl-demo)

Copied to clipboard

Challenge: Concept Over Time Analysis is a machine-learning-based feature that allows users to define, refine, and visualize concepts of interest within an interactive interface.
Approach: They propose to extend the Discourse Analysis Tool Suite with Concept Over Time Analysis extension that allows users to define, refine, and visualize their concepts of interest within an interactive interface.
Outcome: The proposed system allows users to define, refine, and visualize their concepts of interest within an interactive interface.
Reference Free Domain Adaptation for Translation of Noisy Questions with Question Specific Rewards (2023.findings-emnlp)

Copied to clipboard

Challenge: Creating a synthetic parallel corpus from noisy data is also difficult due to its noisy nature.
Approach: They propose a training methodology that fine-tunes the NMT system only using source-side data to balance adequacy and fluency.
Outcome: The proposed method surpasses the MLE-based fine-tuning approach by achieving a 1.9 BLEU improvement.
InVeRo-XL: Making Cross-Lingual Semantic Role Labeling Accessible with Intelligible Verbs and Roles (2021.emnlp-demo)

Copied to clipboard

Challenge: InVeRo-XL is an off-the-shelf system capable of annotating text with predicate sense and semantic role labels from 7 predicated-argument structure inventories in more than 40 languages.
Approach: They propose to use RESTful API and Web interface to integrate sentence-level semantics into cross-lingual downstream tasks.
Outcome: The proposed system can annotate text with predicate sense and semantic role labels from 7 predicated-argument structure inventories in more than 40 languages.
Scaling Multi-Document Event Summarization: Evaluating Compression vs. Full-Text Approaches (2025.naacl-short)

Copied to clipboard

Challenge: Summarizing large text collections is a valuable tool for document research . a multi-stage pipeline and lack of global context are challenges for large-scale summarization systems.
Approach: They compare compression and full-text systems for large-scale multi-document summarization . they find that compression-based methods outperform full-context methods .
Outcome: The proposed methods outperform compression-based methods on three datasets . however, they suffer information loss due to their multi-stage pipeline and lack of global context.
CALE : Concept-Aligned Embeddings for Both Within-Lemma and Inter-Lemma Sense Differentiation (2026.eacl-long)

Copied to clipboard

Challenge: Recent work on Word-in-Context fine-tunes models to investigate lexical meaning but only compares occurrences of the same lemma, limiting the range of captured information.
Approach: They propose an extension to Word-in-Context to include inter-words scenarios by using a dataset and several models on a data set.
Outcome: The proposed models provide efficient multi-purpose representations of lexical meaning that reach best performances in the experiments.
A Multimodal In-Context Tuning Approach for E-Commerce Product Description Generation (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for generating product descriptions from images are inaccurate and generic . e-commerce product descriptions are important for content marketing and increasing engagement .
Approach: They propose a new setting for generating product descriptions from images, augmented by marketing keywords.
Outcome: The proposed approach improves the accuracy and diversity of product descriptions by up to 3.3% on Rouge-L and 9.4% on D-5.
On the Economics of Multilingual Few-shot Learning: Modeling the Cost-Performance Trade-offs of Machine Translated and Manual Data (2022.naacl-main)

Copied to clipboard

Challenge: a framework to evaluate the performance and cost trade-offs between machine-translated and manually-created labelled data is presented.
Approach: They propose a framework to evaluate the performance and cost trade-offs between machine-translated and manually-created labelled data for task-specific fine-tuning of massively multilingual language models.
Outcome: The proposed framework can be used to evaluate cost trade-offs between machine-translated and manually-created labelled data for task-specific fine-tuning of massively multilingual models.
Annotate Chinese Aspect with UMR——a Case Study on the Liitle Prince (2024.lrec-main)

Copied to clipboard

Challenge: Uniform Meaning Representation (UMR) is a graphbased cross-linguistically applicable semantic representation that allows for deep semantic analysis.
Approach: They propose to use an aspectual lattice to adapt to different languages and design values that encompass both viewpoint aspect and situation aspect.
Outcome: The proposed representations are based on the Chinese version of The Little Prince and are compared with other representations.
Self-Adapted Utterance Selection for Suicidal Ideation Detection in Lifeline Conversations (2023.eacl-main)

Copied to clipboard

Challenge: Existing methods for identifying suicidal ideation in phone conversations are difficult to use because of their long duration and noisy nature.
Approach: They propose a self-adaptive approach that identifies the most critical utterances that the NLP model can more easily distinguish.
Outcome: The proposed approach outperforms the baseline models in overall performance with an F score of 66.01% and significantly higher F-score in detecting the most dangerous cases.
MultiHumES: Multilingual Humanitarian Dataset for Extractive Summarization (2021.eacl-main)

Copied to clipboard

Challenge: a new multilingual summarization model is being developed to help humanitarian experts process large amounts of secondary data to derive situational awareness and guide decision-making.
Approach: They propose to use multilingual documents and annotated snippets to improve extraction of secondary data for humanitarian response experts.
Outcome: The proposed model provides multilingual documents with informative snippets that have been annotated by humanitarian analysts over the past four years.
VSCBench: Bridging the Gap in Vision-Language Model Safety Calibration (2025.findings-acl)

Copied to clipboard

Challenge: Existing safety calibration methods focus on model undersafety, where the model responds to hazardous queries, while neglecting oversafetiness, where models refuse to answer safe queries.
Approach: They propose safety calibration which addresses both undersafety and oversafetiness by comparing model responses to a novel dataset of 3,600 image-text pairs.
Outcome: The proposed methods have been used to evaluate safety calibration across image-centric and text-centric scenarios.
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models (2025.findings-naacl)

Copied to clipboard

Challenge: Text-to-image (T2I) models can be used to generate harmful content such as sexually explicit, unfaithful, and misleading or Not-Safe-for-Work (NSFW) images.
Approach: They propose a more practical and universal attack that does not require the presence of a target model.
Outcome: The proposed attack bypasses both text and image safety checkers while preserving high semantic alignment with the target prompt.
LexGLUE: A Benchmark Dataset for Legal Language Understanding in English (2022.acl-long)

Copied to clipboard

Challenge: Laws and their interpretations, legal arguments and agreements are typically expressed in writing.
Approach: They propose a benchmark to evaluate model performance across legal NLU tasks . they also evaluate several generic and legal-oriented models .
Outcome: The proposed model performs better across multiple tasks than previous models.
Understanding the Inner-workings of Language Models Through Representation Dissimilarity (2023.emnlp-main)

Copied to clipboard

Challenge: Dissimilarity measures measure the extent to which two model’s internal representations differ . they can identify and locate generalization properties of models that are invisible via in-distribution test set performance.
Approach: They propose to use representation dissimilarity measures to measure the extent to which two model’s internal representations differ.
Outcome: The proposed dissimilarity measures can identify and locate generalization properties of models that are invisible via in-distribution test set performance and new evaluations of how language model features vary as width and depth are increased.
Detecting Loanwords in Emakhuwa: An Extremely Low-Resource Bantu Language Exhibiting Significant Borrowing from Portuguese (2024.lrec-main)

Copied to clipboard

Challenge: Existing corpora in African languages reveal significant spelling inconsistencies, contributing to poor-quality textual data when encountered in written form.
Approach: They propose a supervised method to identify loanwords in Portuguese . they employ traditional machine learning algorithms incorporating handcrafted features .
Outcome: The proposed method achieves the F1-score of 93% in Emakhuwa, borrowed from Portuguese.
LSC-Eval: A General Framework to Evaluate Methods for Assessing Dimensions of Lexical Semantic Change Using LLM-Generated Synthetic Data (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for measuring Lexical Semantic Change are lacking historical benchmarks.
Approach: They propose a three-stage general-purpose evaluation framework that simulates theory-driven LSC using In-Context Learning and a lexical database.
Outcome: The proposed framework evaluates the sensitivity of computational methods to synthetic change and their suitability for detecting change in specific dimensions and domains.
Can LLMs Understand the Impact of Trauma? Costs and Benefits of LLMs Coding the Interviews of Firearm Violence Survivors (2026.findings-acl)

Copied to clipboard

Challenge: Firearm violence research remains underfunded and difficult to scale due to the lack of funding from the NIH and CDC.
Approach: They use open-source large language models to inductively code interviews with 21 Black men who have survived community firearm violence.
Outcome: The use of open-source LLMs to inductively code interviews with 21 Black men shows that the models can identify important codes, but that they are highly sensitive to data processing.
Enriching a Lexicon of Discourse Connectives with Corpus-based Data (L18-1)

Copied to clipboard

Challenge: Existing annotation efforts for multiple languages have focused on discourse connectives, but we have limited it to the class of connectives marking contrast and the additional relations such connectives might convey.
Approach: They enrich a lexicon of italian COnnectives with real corpus data for connectives marking contrast relations in text.
Outcome: The proposed resource is a valuable tool for linguistic analyses of discourse relations and the training of a classifier for NLP applications.
Good or Bad News? Exploring GPT-4 for Sentiment Analysis for Faroese on a Public News Corpora (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on sentiment analysis in low-resource languages have focused on major languages and emotionally laden text genres like social media and reviews.
Approach: They propose to use GPT-4 for sentiment analysis on Faroese news texts using a multi-class approach with 225 sentences analysed in 170 articles.
Outcome: The proposed model performs remarkably well on 225 sentences and 170 articles compared to human annotators .
Beyond Factual Accuracy: Evaluating Coverage of Diverse Factual Information in Long-form Text Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing evaluation frameworks for large language models focus on isolated aspects like * Equal contribution.
Approach: They evaluate ICAT, an evaluation framework for measuring coverage of diverse factual information in long-form text generation.
Outcome: The evaluation framework is based on three implementations with different assumptions on availability of aspects and alignment method.
From Isolates to Families: Using Neural Networks for Automated Language Affiliation (2025.acl-long)

Copied to clipboard

Challenge: linguistic affiliation of languages to a common language family is traditionally carried out manually . large-scale standardized collections of multilingual wordlists and grammatical language structures could improve this .
Approach: They propose to use lexical and grammatical data to classify languages into families using neural network models.
Outcome: The proposed models outperform models trained on lexical and grammatical data while combining both types of data yields even better performance.
Restoring Ancient Ideograph: A Multimodal Multitask Neural Network Approach (2024.lrec-main)

Copied to clipboard

Challenge: despite efforts to preserve cultural relics, many ancient artefacts have fallen prey to ravages of time, natural deterioration, or deliberate human actions.
Approach: They propose a multimodal multitask restoration model that uses visual and context understanding to restore ancient texts.
Outcome: The proposed model predicts damaged characters and generates restored images simultaneously.
HopWeaver: Cross-Document Synthesis of High-Quality and Authentic Multi-Hop Questions (2026.acl-long)

Copied to clipboard

Challenge: Multi-Hop Question Answering (MHQA) is a critical benchmark for evaluating the model’s ability to integrate information from diverse sources.
Approach: They propose a framework that synthesizes authentic multi-hop questions without manual annotation without the need for manual guidance.
Outcome: The proposed framework synthesizes bridge and comparison questions without human intervention and achieves comparable or superior quality to human-annotated datasets at a lower cost.
Speech Recognition Corpus of the Khinalug Language for Documenting Endangered Languages (2024.lrec-main)

Copied to clipboard

Challenge: Existing tools to document endangered languages are limited due to data scarcity and the need for training.
Approach: They propose to use a speech corpus for Khinalug, an endangered language spoken in northern Azerbaijan, to create a model that can be used in language documentation scenarios.
Outcome: The proposed model achieves 6.65 CER points and 25.53 WER points in low-resource scenarios.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations